Papers with open-ended objective tasks

    1 papers
    F-Eval: Asssessing Fundamental Abilities with Refined Evaluation Methods (2024.acl-long)

    Copied to clipboard

    Challenge: Large language models (LLMs) have been evaluated for their instruction-following capabilities but lack references to their fundamental abilities.
    Approach: They propose a bilingual evaluation benchmark to evaluate the fundamental abilities of large language models including expression, commonsense and logic.
    Outcome: The proposed evaluation methods show higher correlation coefficients and larger distinction than other evaluators.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations